The Plant Phenome Journal
○ Wiley
Preprints posted in the last 90 days, ranked by how well they match The Plant Phenome Journal's content profile, based on 14 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Okyere, F. G. G.; Mehrem, S. L.; Snoek, B. L.; Van den Ackerveken, G.; Abeln, S.
Show abstract
While whole genome sequencing captures millions of single nucleotide polymorphisms (SNPs) and hyperspectral imaging (HSI) enables non destructive plant phenotyping, integrating these modalities to link genotype to phenotype remains challenging due to their high dimensionality and non linearity. This study presents DeepPheno a deep learning framework that predicts SNP genotypes from HSI data, using model predictability as a proxy for genotype phenotype association. HSI data were acquired from 194 lettuce genotypes under field conditions. HSI data patches (20 x 20 pixels x 224 spectral bands) were used to train a hybrid CNN to predict the variant of a specific SNP. The framework was validated on SNPs with known phenotypic effects (anthocyanin, leaf serration, pale pigmentation), achieving high predictive performance (AUC ranging from 0.806 to 0.935), whereas models trained on randomly shuffled labels performed at chance (mean AUC {approx} 0.51). Extending the workflow to 50 randomly selected putatively neutral SNPs, most yielded low predictability, but two showed high performance (AUC > 0.76), suggesting uncharacterized genotype phenotype links. Explainable AI, including SHAP and Grad CAM, identified relevant spectral and spatial features driving these predictions, particularly the green and red edge wavelengths associated with pigment dynamics and leaf structure. These results establish a framework for understanding complex genotype phenotype interactions in plants and extracting these links from HSI data without predefining the exact trait values. It provides an avenue for high throughput trait discovery and description and extends the integration of image based phenomics with plant genetics.
Prusokiene, A.; Prusokas, A.; Retkute, R.
Show abstract
Banana diseases impose severe production losses in tropical smallholder farming systems, yet accurate in-field visual diagnosis remains difficult: symptom expression varies across cultivars and growth stages, and several diseases produce morphologically overlapping foliar signs. We developed a probabilistic image-recognition framework for detecting five economically important banana diseases -- Xanthomonas Wilt, Banana Bunchy Top Disease, Fusarium Wilt (Panama disease), Yellow Sigatoka, and Black Sigatoka -- from in-field photographs, without any disease-specific fine-tuning of the vision backbone. The approach extracts frozen 1,152-dimensional embeddings from the DINOv3 vision foundation model and couples them with a conditional normalizing flow, trained on four publicly available datasets spanning diseased banana plants, healthy tissue, non-banana vegetation, and general natural imagery. On an independent test set the model achieved F1 scores exceeding 0.98, average precision values of 0.968-0.999, and AUROC values of 0.997-1.000 across all five diseases evaluated as binary detection problems. Multi-class accuracy was near-perfect, with limited confusion between Yellow Sigatoka and Black Sigatoka -- a biologically plausible ambiguity attributable to overlapping early-infection foliar symptoms. Because the normalizing flow estimates explicit conditional probability densities rather than decision boundaries, two complementary log-likelihood ratios can be derived: a disease ratio comparing each disease class against healthy banana, and a plant ratio comparing banana against non-banana imagery. Together these define an interpretable two-dimensional diagnostic space that simultaneously quantifies evidence for disease presence and image relevance, cleanly separating diseased plants, healthy plants, and out-of-distribution images while flagging uncertain predictions for confirmatory testing. Inference on frozen embeddings is lightweight and compatible with smartphone deployment, providing a scalable, uncertainty-aware diagnostic tool for smallholder farming systems and disease surveillance programmes.
Nguyen, T. V.; Quoc, K. N.; Harwath, D.; Quach, L.-D.; Dao, P. D.
Show abstract
Plant diseases remain a major challenge to global food production, and timely, accurate, and scalable detection of plant stress is critical to reducing these losses. Recent advances in digital imaging and artificial intelligence offer unprecedented opportunities for precision crop disease detection and management. Yet, existing plant disease datasets remain often fragmented across crop and disease systems, and are largely dominated by controlled-environment imagery. The lack of standardized, interoperable, and representative datasets limits reproducibility, transferability, and scalability of AI systems, thereby constraining their deployment in operational agricultural applications. Here we present LeafMD, an integrated multimodal plant disease dataset and benchmark resource that includes LeafNet 2.0, a large-scale multimodal digital image dataset comprising 255,855 image-text pairs across 37 crop species, 197 crop-disease classes, and 9 geographic regions spanning tropical, subtropical, and temperate agricultural systems. Unlike conventional datasets, LeafNet 2.0 integrates biologically grounded symptom descriptions with image-level annotations of early and late disease stages, enabling symptom-aware analysis of disease progression under realistic field conditions. We further introduce LeafBench 2.0 as part of LeafMD, a visual-question answering benchmark covering nine fine-grained plant pathology tasks, including pathogen classification, lesion characterization, symptom interpretation, and disease severity assessment. Evaluation across 16 vision-language models revealed substantial performance gaps between coarse disease recognition and fine-grained pathological reasoning, while agriculture-adapted models consistently outperformed several larger general-domain architectures on symptom-oriented tasks. Together, LeafNet 2.0 and LeafBench 2.0 establish LeafMD as a multimodal resource for developing disease-aware agricultural foundation models and studying fine-grained pathological reasoning in real-world environments.
Matuszynska, A.; Sansa, O.; Adekoya, F. J.; Akinyemi, O. O.; Anokye, E.; Bashir, O. B.; Boyny, Z. Z. F.; Chukwuka, M. K.; Corvest, E.; Dada, A. O.; DellAcqua, M.; Ehemba, G. L.; Finkbeiner, A. J.; Hamabwe, S.; Hodehou, D. A. T.; Kacheyo, O.; Kamfwa, K.; Mhango, K. J.; Abdullahi, W. M.; Munduwe, G.; Ntukidem, S.; Obisesan, O. K.; Odesina, I. S.; Ogechi, N.-U.; Olaoye, O. D.; Olayinka, M. M.; Osei-Bonsu, I.; Rilwan, K. O.; Stival, L.; Tehar, Z.; Tende, R. M.; To, J.; Ugochukwu, U. K.; Unger, A.; van Aalst, M.; Vrbic, D.; Zhang, C.; Theeuwen, T. P. J. M.; Kramer, D. M.; Kromdijk, J.
Show abstract
Photosynthesis is among the most consequential yet genetically complex traits in crop plants, and translating its natural variation into actionable genomic targets remains a central challenge for breeding climate-resilient varieties. To start addressing this, researchers are generating increasingly large, multi-environment field photosynthesis datasets. Yet, these data have been structurally under-analysed since their inception. Here we report the outcomes of the first dedicated hackathon focused on computational mining of such field data held in Accra, Ghana, in March 2026. Bringing together data scientists, plant physiologists, geneticists, and breeders from Europe and Africa, these interdisciplinary teams used photosynthetic data collected with hand-held fluorometers to genome-wide marker data across four crop species: cowpea (Vigna unguiculata), barley (Hordeum vulgare), common bean (Phaseolus vulgaris), and potato (Solanum tuberosum). Despite using different species and methods, independent teams identified the same three key findings. First, mechanism-informed feature engineering and dynamic modelling recover genetic signals that are not detected or discarded in standard analysis pipelines, resulting in traits with improved heritability and meaningful associations with yield. Secondly, machine learning methods proved effective at uncovering genetic associations, with temporally resolved features substantially outperforming single time-point measurements. Third, raw chlorophyll fluorescence and absorbance traces consistently contained more information and predictive power than the extracted parameters currently used. A defining feature of this event was having experimentalists and data scientists working together, enabling AI approaches to be grounded in domain knowledge and biological mechanisms rather than relying on data alone.
Parth, K.; Varela, S.; Liu, Z.; Martini, K. M.; Rajurkar, A.; Allan, D.; McCoy, S.; Ruhter, J.; Walker, S.; Goldenfeld, N.; Leakey, A.
Show abstract
Quantifying root traits such as root length (RL) and root surface area (RSA) from minirhizotron imagery is a valuable approach for overcoming the phenotyping bottleneck that limits understanding and improvement of crop productivity, resource use efficiency and resilience in field experiments. However, current approaches remain labor-intensive, and deep learning (DL) methods suffer from limited generalization ability. We present RootQuant, an end-to-end DL model that simultaneously predicts RL and RSA directly from minirhizotron images using only whole-image trait values as supervision, thereby eliminating the need for pixel-level annotations. The models generalization ability was evaluated across species and fine-tuning configurations. The practical applicability of the model was further assessed under field conditions by converting image-derived RL estimates into volumetric root length density (vRLD). Using 118,191 maize and soybean images collected between 2009 and 2020, RootQuant trained on both species achieved an R2 of 0.90 and an RMSE of 2.9 mm for RL, and an R2 of 0.88 and an RMSE of 4.2 mm2 for RSA. The same mixed-species model generalized strongly across species, yielding an 8% relative improvement in R2 and a 30% lower RMSE on maize compared with the same architecture trained on a single species and applied zero-shot. Image-derived RL predictions converted to vRLD showed the expected depth-dependent decline in vRLD, as was also found by coincident destructive quantification of roots washed out of soil cores. By providing a generalist backbone model trained on a large dataset from two major crop species, RootQuant enables high-throughput simultaneous estimation of two relevant root traits directly from raw imagery without task-specific fine-tuning, thereby accelerating in situ root system analysis and phenotyping applications.
Harris, Z. N.; Braley, J.; Cassetta, E.; Crain, J.; DeHaan, L.; Van Tassel, D.; Miller, A.; Rubin, M. J.
Show abstract
Perennial grains represent a promising frontier for sustainable agriculture, but breeding progress is constrained by the accessibility of genotyping and the difficulty of evaluating complex traits expressed for multiple years after establishment across heterogeneous environments. Phenomic selection may help address these challenges by using inexpensive, scalable, high-dimensional phenotypes collected early in development, although the robustness of such predictions across breeding cycles remains uncertain. Here, we compared genomic selection and phenomic selection across two breeding cycles of Thinopyrum intermedium (intermediate wheatgrass; IWG; Kernza(R)), comprising approximately 2,280 individuals from maternal half-sib families evaluated across multiple field sites and years. We constructed relationship matrices from genomic markers and early-life stage phenomic data, including seed and leaf color (HSV), CropReporter multispectral reflectance and indices, and cycle-specific hyperspectral reflectance sensors. Genomic models provided the strongest predictions on average across all field traits in both cycles. Among phenomic predictors, leaf HSV was consistently the most informative, whereas CropReporter and hyperspectral data showed lower and more trait-dependent performance and seed HSV provided little predictive value. Genomic, leaf HSV, and CropReporter models transferred across breeding cycles with little apparent loss of predictive ability relative to within-cycle validation, demonstrating that their predictive signals were not restricted to a single breeding cycle. Early-life stage leaf HSV emerged as a practical, accessible tool for germplasm thinning and early-stage prioritization in perennial breeding programs. Despite limited similarity among relationship matrices, multi-relationship-matrix models rarely improved prediction beyond the stronger constituent single-relationship-matrix model. Together, these results show that early-life stage phenomic data provide reproducible information about agronomic performance expressed years later, but that predictor complexity and data integration do not guarantee improved prediction.
Stock, F.; Panda, S.; Poire, R.; Brown, T.; Akram, A.; Zheng, L.; Lei, H.; Zha, R.; Zhao, M.; Isabelle, S.; Martel, M.; Comeau, M.-A.; Hamel, L.-P.; Lavoie, P.-O.; D'Aoust, M. A.; Reithinger, H.; Saxena, P.; Stone, E. A.; Li, H.; Way, D. A.; Atkin, O. K.
Show abstract
Non-invasive, high-throughput phenotyping tools are needed that can identify environmental effects on plant structure and function to diagnose factors responsible for reduced growth in commercial and non-commercial settings. In this study, we explored whether the integration of 3D-multispectral (3D) and 2D-hyperspectral imaging (HSI), aided by machine learning (ML), could be used to identify environmental stress treatments imposed during plant growth. Controlled environment-grown Nicotiana Benthamiana plants were subjected to a range of abiotic treatments - including different growth irradiances, heat treatment and drought stress - with the treatments resulting in differences in shoot height, biomass, leaf area and spectral reflectance. ML models were trained to identify these treatments using morphological and spectral traits measured at 27, 29, 31, and 34 days after sowing (DAS). A 3D-multispectral scanner was used to obtain information on plant height, biomass, and leaf area. A visible and near-infrared (VNIR) HSI camera provided detailed spectral information for deriving spectral indices including the Normalised Difference Vegetation Index (NDVI), Photochemical Reflectance Index (PRI) and Normalized Difference Red Edge (NDRE). Manual measurements provided baseline comparative data. The 3D-multispectral scanner reliably estimated above-ground traits, with high correlations between manual and scanner-derived measurements. The ML models accurately differentiated among environmental stress treatments, with the fused 3D+HSI model achieving the best overall predictive performance across all evaluated metrics compared with models based on either imaging modality alone. Results demonstrated the effectiveness of combining 3D-multispectral and 2D-HSI data with ML analyses for non-destructive, high-throughput phenotyping. The integration of these techniques enabled non-destructive, high-throughput identification of environmental stress treatments imposed during plant growth.
Varela, S.; Ruhter, J.; Sacks, E.; Zheng, X.; Allen, D.; Hale, A.; Landry, C.; Kuang, X.; Long, B.; Zhu, Y.; Proma, S.; Kaur, S.; Jarquin, D.; Morrison, J.; Leakey, A.
Show abstract
The integration of digital technologies for high-throughput field phenotyping is critical for accelerating crop improvement in agriculture. However, extracting traits from remote sensing data remains constrained by fragmented workflows, manual intervention, and limited interoperability among existing tools, resulting in delays that hinder timely biological insight and decision-making. To address these challenges, we present PhenoStream (Phenotyping Streaming), a scalable, end-to-end cyberinfrastructure designed to automate the full lifecycle of aerial imagery-based phenotyping, from data acquisition to plot- and genotype-level inference. The framework integrates automated data ingestion from distributed field sites, geospatial processing, and AI-enabled trait extraction within a unified, user-accessible graphical interface. Its modular and extensible architecture supports adaptable trait modeling and seamless integration of new data sources, enabling deployment across diverse crops, environments, and experimental designs. We demonstrate the system across a large multi-location field trial network of bioenergy crops, where it enables high-throughput characterization of spatiotemporal growth dynamics, genotype-by-environment (GxE) interactions, and predictive modeling of key agronomic traits. By significantly reducing processing latency and manual effort, the platform facilitates near-real-time analysis and reproducible workflows. This work establishes a generalizable and scalable pathway for operationalizing very-high-spatial resolution aerial phenotyping in agricultural research. By bridging data acquisition and analytics, the end-to-end cyberinfrastructure provides a foundation for integrating heterogeneous and unstructured data streams--including remote sensing, environmental, and management data--toward data-driven decision making in agriculture.
Sims, B.;Gaudinier, A.;Blackman, B.
Show abstract
PremiseSeed size and morphology are critical traits in agriculture, ecology, and genetics, but high-throughput quantification of these traits is often limited by labor-intensive manual measurements or expensive, platform-specific imaging software. Methods and ResultsWe developed SeedMeasure, a lightweight, open-source, and cross-platform command-line tool written in Python that automates the measurement of seed area, length, and width from images. Using a simple imaging setup, the program processes images by correcting for perspective skew, filtering debris, and exports quantitative data alongside quality-check images. We validated SeedMeasure across nine diverse species, ranging from small Arabidopsis thaliana seeds to large Zea mays kernels. The tool quickly handles images using multithreading and demonstrates high reproducibility, yielding low coefficients of variation across repeated runs. ConclusionsCompared to existing software, SeedMeasure is free, offers faster processing through parallel computing, and provides standalone executables that require no programming dependencies. SeedMeasure offers an accessible, cost-effective, and high-throughput approach for rapid phenotypic profiling, making advanced seed morphological analysis available to researchers without specialized laboratory hardware.
Mandelli, L.; Johnson, K. M.; Berretti, S.; Mencuccini, M.
Show abstract
1O_LIEmbolism, the formation of air bubbles in the plant water transport system, is a mechanistic driver of plant death. The Optical Vulnerability Technique (OVT) is an imaging method for non-invasive quantification of embolism (including P50, a common metric for drought vulnerability), which can also provide detailed spatial and temporal information. Its major cost lies in the post-processing of thousands of images. C_LIO_LIHere we designed, tested, trained, and make publicly available a neural network model to automate post-processing of OVT images. Using a dataset of 65 leaves from Senecio pterophorous, we compared our model predictions to results obtained via traditional post-processing by an expert. C_LIO_LIOur model resolved P50 to within 0.027 MPa of the expert-processed data with training taking 30 minutes to 2.5 hours and model-runtime in the order of seconds to minutes, demonstrating its promise for increasing the efficiency and throughput of P50 calculation. The models performance in replicating the pixels that constitute embolism events was lower (mean event-frame IoU of 0.38). C_LIO_LIWe invite the community to utilise our model but emphasise that it does not replace the expert-processing pipeline and that care must be taken when considering applying this and similar approaches to OVT data. C_LI
Dubois, R.; Bousset, L.; Jumel, S.; Leclerc, M.; Parisey, N.; Joly, A.
Show abstract
Accurate segmentation of plant disease symptoms is essential for crop monitoring and phenotyping, yet it typically requires costly pixel-level annotations. Weakly supervised semantic segmentation (WSSS) alleviates this burden using image-level labels, but its performance depends on the quality of spatial priors such as class activation maps (CAMs). We investigate whether text-guided segmentation with the Segment Anything Model 3 (SAM3) can serve as an alternative weak supervision signal. Three pseudo-mask generation strategies are compared: (i) CAMs refined with SAM or SAM3, (ii) zero-shot text-guided SAM3, and (iii) a hybrid approach combining weak spatial cues with text prompts. The resulting pseudo-masks are used to train a DeepLabV3 model. Text guidance alone matches or outperforms conventional WSSS, achieving up to 0.46 IoU without spatial supervision and 0.61 IoU on a public dataset, although performance is sensitive to text prompt formulation. The hybrid strategy improves robustness, reaching 0.50 IoU on the primary dataset and 0.58 IoU on the additional dataset while reducing prompt sensitivity. Overall, text guidance is a promising alternative to conventional weak supervision, while hybrid approaches provide a more robust solution for plant disease segmentation.
Okyere, F. G. G.; Mehrem, S. L.; Snoek, B. L.; Van den Ackerveken, G.; Abeln, S.
Show abstract
Understanding the link between genetic variation and observable traits is key to crop breeding. Hyperspectral imaging captures physiological and biochemical profiles, but current supervised methods require costly trait annotations and treat each observation as a static snapshot, ignoring the temporal dynamics of plant development. We introduce SST-MAE, a self-supervised framework that learns genotype-discriminative representations from plant hyperspectral developmental trajectories, without requiring phenotypic labels. The model learns to reconstruct masked information, capturing multiple growth trajectories. Validated on 194 field-grown lettuce genotypes across eight time points, the frozen encoder serves as a feature extractor for downstream genotype classification. SST-MAE outperforms raw spectral and linear baselines, achieving AUROC > 0.89 for anthocyanin pigmentation SNPs and 0.77 for leaf serration. The learned features are highly label-efficient, attaining near-full performance with only 30-50% of labeled data, offering a scalable pathway toward high-throughput genetic screening from image-based phenotypes.
Daware, A. v.; Hacke, C.; Remay, A.; Starnberger, P.; Schraml, C.; Collonnier, C.; Laurens, F.; Schmid, K. J.
Show abstract
Testing for distinctness, uniformity, and stability (DUS) is a requirement for plant variety registration and based on phenotypic traits, which is time-consuming and sensitive to environmental variation. Advances in genomics allow to complement DUS testing with molecular markers, for which two models in DUS testing were proposed by the Union for the Protection of New Varieties of Plants (UPOV). A use cases was described for maize, but an implementation has been hindered by a lack of suitable markers and validated analytical frameworks. We address these challenges by integrating historical DUS characteristics scores from 352 European hybrid maize varieties with high-density genome-wide single nucleotide polymorphism (SNP) data. Using genome-wide association studies (GWAS), we identified 18 genomic regions and candidate genes associated with 12 DUS characteristics, enabling the development of diagnostic markers consistent with the UPOV model "Characteristic-Specific Molecular Markers". Since most DUS traits are polygenic, we combined GWAS-informed marker selection with XG-Boost-based machine learning to predict notes of DUS characteristics. This approach achieved strong predictive performance across multiple traits (mean accuracy 0.67), demonstrating its potential for managing reference collections under UPOV model "Combining phenotypic and molecular distances in the management of variety collections". Both approaches were validated for two characteristics using independent public USDA-NPGS maize datasets (>1,700 accessions) highlighting the value of public data for method validation. We also identify key limitations of historical DUS data, including imbalanced and sparse trait representation, and discuss mitigation strategies. Despite these constraints, our results demonstrate that molecular markers may improve maize DUS testing, enabling faster, more accurate variety registration and supporting accelerated crop improvement. Key messageHistorical DUS datasets can be used to identify marker-trait associations of DUS characteristics using genome-wide association study (GWAS) and to develop a genomic prediction framework for an accurate prediction of DUS character notes from marker data.
Feng, Q.; Rafter, P.; Wilson, I.; Li, Z.; Conaty, W.
Show abstract
ContextUnderstanding how and when environmental conditions influence overall crop performance is crucial for optimising the development of genotypes to a specific breeding target environment. We focused on economically important traits of Australian rain-grown cotton including fibre yield and quality traits, which have not been investigated comprehensively. The aim of the study was to identify relevant environmental factors, and the timing and extent of their impact on rain-grown cotton production. MethodsWe used a data driven approach to analyse the relationship between ten climate related environmental factors across various plant growth stages and eight fibre yield and quality traits, using a large-scale field dataset of 9,283 records collected over 23 years at 4 locations, with 53 unique year-location combinations. We applied eight complementary statistical models including stepwise, penalised and Bayesian linear regression, regression-tree based ensemble methods and deep learning frameworks to (1) select the most essential environmental covariates affecting rain-grown cotton production, and (2) evaluate the predictive performance of these models. ResultsThe environmental impacts on rain-grown cotton production were trait and growth-stage specific. Number of rainy days and solar radiation were identified as the most influential environmental factors for fibre yield traits, vapour pressure deficit at maximum daily temperature was the most influential factor for majority of fibre quality traits. However, each analysed trait was influenced by multiple environmental factors across multiple growth stages (rather than a single factor or a single growth stage). These influential covariates explained a wide range of variation in the traits, accounting for 5.8% to 68.2%. Using the best-fit random forest model, our findings revealed non-linear relationships between key environmental covariates and the traits. ConclusionsEnvironmental factors at different rain-grown cotton growth stages are key determinants for the performance of end-of-season fibre yield and fibre quality parameters. These findings highlight the need to account for environment conditions when developing cotton varieties optimised for rain-grown production systems. Potential strategies are proposed whereby these key environmental factors can be used to increase the rate of genetic gain in rain-grown cotton production systems. ImplicationsThe results of this study will be crucial for future genetic evaluations and analyses of genotype-by-environment interaction effects in rain-grown cotton, which must account for the influence of the environment on plant performance. Furthermore, these methods can be applied to other species to identify critical growth stages and environmental factors which most influence crop performance.
Mejias, J.; Adreit, H.; Blanc, A.; Lubin, N.; Jolivet, C.; Guyot, V.; Brayle, O.; Poncelet, N.; Fournier, E.; Wicker, E. P.; Carlier, J.; Tharreau, D.; Ravel, S.
Show abstract
BackgroundThe quantification of fungal spores constitutes a fundamental metric in phytopathology, serving as the primary variable for inoculum standardization and being used as a proxy for disease severity. Historically, spore quantification has relied on manual hemocytometry, which remains the most precise counting process to date, where chambers such as the Malassez slide are used to count a subsample of the inoculum. However, this method applied manually is highly labor-intensive, time-consuming, and can be prone to operator-dependent variability. To overcome these limitations, we introduce MIRA (Microscopy Image Recognition & Analysis), a novel open-source software integrating You Only Look Once (YOLO) deep learning algorithms. Featuring a user-friendly graphical interface, MIRA is adaptable to multiple camera systems and supports advanced object detection models, including YOLOv11 and YOLOv26. ResultsWe demonstrate that MIRA can be used to accurately detect and count spores from several phytopathogenic fungi, automatically measure spore surface area, and to differentiate spores across different genera. In an exhaustive comparative analysis using Pyricularia oryzae spores as an example, MIRA was benchmarked against manual gold-standard counting slides (Malassez and Kova) and indirect spectrophotometric methods (SPARK). The P. oryzae model loaded via MIRA achieved a strong correlation (R = 0.96) with manual gold standards while reducing processing time by over 90% for high-concentration samples (10 spores/mL). Beyond this benchmark, we also successfully tested specific YOLO models designed to recognize macro- and microconidia of Fusarium oxysporum f. sp. cubense, a model for Pseudocercospora fijiensis, and a single multiclass model capable of identifying six different rice pathogenic fungi. We provide comprehensive tutorials for operating the software and training custom detection models for free using Roboflow and Google Colab. MIRA is available both as open-source Python code and as standalone executables for Windows and Linux. ConclusionsMIRA provides a rapid, accurate, and highly reproducible alternative to manual spore counting, effectively removing a major bottleneck in phytopathology workflows. By combining advanced YOLO-based deep learning with an accessible interface and comprehensive training resources, MIRA makes accessible automated image analysis for researchers without programming expertise. Moreover, MIRA drastically improves the efficiency of high-throughput disease phenotyping and can be adapted for a wide range of microscopic quantification tasks across various biological disciplines.
Severini, A. D.; Gawinowski, M.; Bancal, M.-O.; Launay, M.; Deswarte, J.-C.; Chenu, K.
Show abstract
Crop models are essential for predicting climate change impacts on agriculture, yet their validation under multi-stress conditions remains limited. This study evaluated two widely-used wheat models, APSIM and STICS, using data from three Free-Air CO2 Enrichment (FACE) experiments (USA, Germany, Australia) combining elevated CO2 (eCO2), water deficit, and warming. Environmental characterisation using simulation-based stress indices revealed that intended "controls" frequently experienced hidden heat and water stress, meaning models were calibrated on crops already undergoing physiological adjustments. Evaluation of simulated yield and components revealed a clear hierarchy in prediction errors (RRMSE): unlimited conditions (3-9%) < single stress (4-27%, with a need to improve response to heat stress) < combined stress (17-123%). Elevated CO2 generally increased prediction uncertainty for crops experiencing water stress. Our results suggest that current stress functions from the models fail to capture the synergistic coupling between drought and heat stress. This highlights the urgent need for more mechanistic modelling to improve the reliability of climate change impact assessments.
Nakata, R.; Hiraga, S.; Ishimoto, M.
Show abstract
Background and aims Plant volatile organic compounds (VOCs) change dynamically with plant development and in response to environmental conditions. However, their potential as non-invasive indicators of phenological progression remains poorly explored. In this study, we developed a framework integrating automated VOC sampling, time-resolved VOC profiling, and machine-learning analysis for the non-invasive assessment of plant phenology. Using soybean (Glycine max (L.) Merr.), we investigated whether development-associated temporal variation in VOC emissions could delineate and predict developmental phases. Methods We collected VOCs daily under controlled environmental conditions from 16 to 43 days after sowing, spanning the transition from vegetative to reproductive stages, using an automated sampling system coupled with thermal desorption-gas chromatograph-mass spectrometer (TD-GC-MS). To characterise temporal changes in VOC profiles associated with phenological progression, we analysed the daily VOC data using a multi-step pipeline combining statistical filtering and similarity-based network analysis. We defined VOC-derived developmental phases from similarity patterns in the VOC profiles, then developed and evaluated machine-learning models to predict these phases. Key results Seven VOCs exhibited distinct phase-dependent dynamics, including green leaf volatiles and monoterpenes showing characteristic temporal changes during phenological progression. Network-based clustering of VOC profiles resolved five developmental phases closely aligned with conventional developmental stages. A machine-learning model predicted these phases from the VOC profiles with high predictive accuracy on independent test data, demonstrating that phenological progression could be quantitatively inferred from VOC emission patterns. Conclusions Our findings support VOC profiling as a reliable and non-invasive approach for assessing phenological progression in soybean. By extracting temporally structured VOC signals, this framework captures developmental information that may be difficult to obtain through visual observation alone, particularly after canopy closure. VOC profiling offers a practical tool for monitoring crop developmental dynamics and has broader potential for plant phenotyping and precision crop management.
Check, J. C.; Technow, F.; Totir, L. R.
Show abstract
Characterizing hybrid maize disease resistance is a costly and labor-intensive effort in commercial breeding programs. Field trials are carefully inoculated and managed but remain error-prone due to spatial variability in disease pressure, microclimatic conditions and inter-rater variability. Quantitative ordinal disease rating scales are used to increase scoring speed at the expense of resolution, accuracy, and the ability to use conventional statistical methods. To improve traditional methods of disease resistance characterization, we propose to leverage readily available low-density SNP marker profiles to create genome-informed disease scores. Specifically, a whole genome ordered probit regression (WGOPR) model is used to deconstruct field-observed disease phenotypes into marker effects and reconstruct genome-informed disease scores. This approach is demonstrated in hybrid maize using data from Exserohilum turcicum-inoculated field trials across the central and northern U.S. and Canadian Corn Belt in 2024. Resulting Genomic Estimated Categorical Probabilities (GECPs) are compared to observed frequencies of disease scores to validate the methodology and evaluate the accuracy of regional hybrid maize disease resistance characterization. The benefit of a probabilistic output is demonstrated through two use cases: a comparison of hybrids with highly variable observed disease resistance scores at a single location, and a comparison of breeding selection schemes from a regional analysis. Because GECPs are the product of estimated marker effects, they better represent the expected behavior of a genotype independent of location-, rater- and plot-specific noise, and will therefore offer a step towards improving hybrid maize characterization and better informing breeding decisions.
Castillo, M. P.; Oyebode, O. G.; Lenahan, A.; Orloski, A.; Wolfe, M.
Show abstract
White lupin (Lupinus albus L.) is a cool-season grain legume with seed crude protein of 33-47%, competitive with soybean (Glycine max L.) meal. It also fixes nitrogen and mobilizes soil phosphorus. Because soybean is a summer crop, white lupin can occupy Southeastern winter fields as a complementary protein source. Breeding for seed protein is limited by the cost and throughput of reference phenotyping. To determine how each is best deployed, we compared the utility of near-infrared spectroscopy (NIRS)-based phenomic selection with genomic selection based on 246,847 SNPs from low-pass, whole genome sequencing in a panel of Auburn University breeding lines and USDA National Plant Germplasm System germplasm. A handheld NIR calibration against Dumas reference protein reached screening-grade accuracy (R2 = 0.81). Under common cross-validation, phenomic predictive ability was 0.93 and genomic was 0.12. The low genomic value was consistent with moderate heritability (H2 = 0.33) and strong genotype-by-year interaction. Beyond predictive ability, NIRS recovered superior accessions the strictest selection intensity, and 40 to 60 reference assays sufficed to calibrate the model. Handheld NIRS is a low-cost tool for protein calibration and early-generation screening, while genomic prediction remains suited to parental selection, together supporting a complementary strategy for legume breeding Plain Language SummarySoybean meal is the main protein source for livestock and fish farms in the United States. Because soybean is a summer crop, many Southeastern fields sit idle or grow low-value cover crops in winter. White lupin, a cool-season legume whose seeds are as protein-rich as soybean meal, makes a good complementary winter crop: it yields high-protein grain while serving as a cover crop that fixes nitrogen and frees up soil phosphorus for later crops. In our early-stage lupin breeding program, measuring seed protein by standard lab methods is slow and costly. We built a calibration that lets a handheld scanner estimate protein from light, and compared it with predicting protein from the plants DNA. The scanner gave accurate, low-cost protein screening from only about 40-60 lab tests, while DNA-based prediction remains suited to guiding parent selection. Used together, these tools offer breeders a practical path to develop high-protein white lupin. Core ideasO_LIHandheld NIRS provides screening-grade prediction of white lupin seed crude protein. C_LIO_LISpectra carried more usable protein signal than markers by measuring seed chemistry directly. C_LIO_LINIRS and genomic prediction serve different stages of a white lupin breeding program. C_LIO_LIAbout 40 to 60 reference assays sufficed to calibrate NIRS to near-full accuracy. C_LI
Sharma, S.; Gustin, J. L.; Frei, U. K.; Settles, A. M.; Lübberstedt, T.; Resende, M. F. R.; Hershberger, J.
Show abstract
Key messageA single-kernel near-infrared reflectance spectroscopy-based sorter can effectively identify haploid kernels for doubled haploid production in field and sweet corn backgrounds. Doubled haploid (DH) technology significantly shortens the breeding cycle for developing homozygous inbred lines in maize (Zea mays). Manual sorting of haploids from a larger bulk of hybrid kernels in an induction cross is a major bottleneck in DH development. Automated systems based on near-infrared (NIR) reflectance spectroscopy can be valuable tools for rapid haploid sorting, provided that sorting accuracy is sufficient for incorporation into the DH process. In this study, we evaluated the accuracy of a custom-built single-kernel NIR (skNIR) sorter for classifying haploid kernels from 12 high-oil haploid induction populations generated from two sweet corn and two field corn donors and four high-oil haploid inducers (HOHIs). We evaluated several general classification models that can be applied without population-specific recalibration or prior genotyping, including models that classified haploids based solely on predicted oil content, as well as multivariate methods that used all wavelengths of the NIR spectra. The highest classification accuracy was obtained using a general multivariate support vector machine (SVM) model. When combined with the two best-performing HOHIs, the general SVM model accurately sorted induction populations from two of the three donor backgrounds crossed with these inducers. Two oil-based methods showed less accurate classification than the multivariate SVM model, due to overlapping oil content distributions across the two kernel classes. Overall, this study demonstrates effective skNIR-based sorting of haploid kernels from diverse induction populations using a single general model. The practical deployment of this instrument in maize breeding programs is discussed.